Parallel Characteristic Extraction from Protein Sequence Database
نویسندگان
چکیده
An adaptive massively parallel system for flexible information processing has been investigated. This research requires a feedback from the real application. In this paper, a parallel characteristic extraction from the protein sequence database is described. Since the protein sequence database is huge and sequences have variety, an adaptive massively parallel system is mandatory. An HMM (hidden Markov model) is essentially suitable for the massively parallel system. An HMM also represents conditional probabilities to deal with the stochastic nature of the protein. However, finding the optimal HMM topology for protein is a hard problem. Thus, an iterative duplication method is developed for HMM topology learning. Using this method, a motif, one of the protein characteristics, is extracted. We obtained an HMM for a leucine zipper motif. Comparing to the accuracy of a symbolic representation which accuracy is 14.8 %, an HMM achieved 79.3 % in prediction. We demonstrated that this approach is applicable to the validation of the protein database; a constructed HMM has indicated that one protein sequence annotated as a “leucine-zipper like” in the database is quite different from other leucine-zipper sequences in terms of likelihood, and we found this discrimination is plausible. 2 HMMs
منابع مشابه
iProsite: an improved prosite database achieved by replacing ambiguous positions with more informative representations
PROSITE database contains a set of entries corresponding to protein families, which are used to identify the family of a protein from its sequence. Although patterns and profiles are developed to be very selective, each may have false positive or negative hits. Considering false positives as items that reduce the selectiveness of a pattern, then, the more selective pattern we have, a more accur...
متن کاملInvestigation of Consecutive Separating Arrangements of Bio active Compounds from Black Tea (Camellia sinensis) Residue
Every year lots of black tea (Camellia sinensis (L.) Kuntze) residue will produce in the factories. These residue are unusable whereas the bio active compounds can be extracted and used in the drag and food industries. Due to mentioned problems, this project was conducted years 2011 - 2012 with the aim to make a study on consecutive isolation of all bio active compounds from tea residu...
متن کاملHigh-throughput comprehensive analysis of human plasma proteins: a step toward population proteomics.
A high-throughput (HT) comprehensive analysis approach was developed for assaying proteins directly from human plasma. Proteins were selectively retrieved, by utilizing antibodies immobilized within affinity pipet tips, and eluted onto enzymatically active mass spectrometer targets for subsequent digestion and structural characterization. Several parameters, including uniform parallel protein e...
متن کاملThe SYSTERS Protein Family Database in 2005
The SYSTERS project aims to provide a meaningful partitioning of the whole protein sequence space by a fully automatic procedure. A refined two-step algorithm assigns each protein to a family and a superfamily. The sequence data underlying SYSTERS release 4 now comprise several protein sequence databases derived from completely sequenced genomes (ENSEMBL, TAIR, SGD and GeneDB), in addition to t...
متن کاملMolecular characterization of apolipoprotein A-I from the skin mucosa of Cyprinus carpio
Apolipoprotein A-I is the most abundant protein in Cyprinus carpio plasma that plays an important role in lipid transport and protection of the skin by means of its antimicrobial activity. A 527 bp cDNA fragment encoding C terminus part of apoA-I from the skin mucosa of common carp was isolated using RT-PCR. After GenBank database searching, a partial sequence containing a coding sequence (CDS)...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 1994